Papers with Large Audio Language Models

9 papers
Audio Query Handling System with Integrated Expert Models and Contextual Understanding (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing chatbots are limited to specific audio tasks, but the domain of audio content related queries remains underexplored.
Approach: They propose to use an intent classifier to route queries to audio-related experts using a diverse audio query dataset.
Outcome: The proposed system outperforms state-of-the-art LLMs on custom audio tasks and MMAU sound set benchmarks.
Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Recent literature focuses on constructing large audio language models (LALMs) but they are limited in temporal reasoning, which may hinder commercial applications .
Approach: They propose a data augmentation technique for generating reliable audio temporal questions and answers using an LLM.
Outcome: The proposed model performs well on public audio benchmark datasets and is optimized for edge applications.
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)

Copied to clipboard

Challenge: Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency.
Approach: They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category .
Outcome: The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category .
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large Audio Language Models (LALMs) have demonstrated unprecedented capabilities in natural language understanding and generation, revolutionizing human-machine dialogue.
Approach: They propose an unsupervised safety-fine-tuning strategy that reshapes LALMs representation space to enhance existing LALM safety-alignment while balancing the risk of over-rejection.
Outcome: The proposed approach improves LALMs safety under three input conditions while increasing over-rejection rate by only 0.88% on average.
SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on noise lack quantitative analysis and rely on intuition and empirical observation, thus failing to understand practical robustness.
Approach: They propose a method for quantifying the impact of noise intensity on LALM inputs by using a structured activation subspace derived from the model's internal representations.
Outcome: The proposed method outperforms existing denoising methods and demonstrates that noise is perceived more accurately than raw audio features.
CORD: Bridging the Audio–Text Reasoning Gap via Weighted On-policy Cross-modal Distillation (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models (LALMs) exhibit a degradation in knowledge and reasoning capabilities . empirical results show that CORD significantly bridges the audio–text performance gap .
Approach: They propose a framework that performs online cross-modal self-distillation to bridge the acoustic-semantic gap between LALMs and text-based models.
Outcome: The proposed framework bridges the acoustic-semantic gap between LALMs and text-based models . it employs on-policy reverse KL divergence with importance-aware weighting .
Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to determine whether to perform reasoning lack fine-grained mechanisms to adapt reasoning length to problem complexity.
Approach: They propose a difficulty-adaptive reasoning method that dynamically links reasoning length to the model’s perceived problem difficulty.
Outcome: The proposed method reduces average reasoning length by 50%, achieving higher efficiency without sacrificing accuracy.
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Recent Large Audio Language Models (LALMs) have shown strong capabilities in audio understanding, yet their reasoning remains vulnerable to perceptual errors.
Approach: They propose a large-scale dataset for **Perception-Aware Question Answering** that uses a hierarchical decoupling strategy to separate speech from environmental sounds and distinguishes among multiple speakers.
Outcome: The proposed model improves on MMAU-mini, MMAR, and PAQA while maintaining comparable performance on multiple benchmarks.
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited .
Approach: They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training.
Outcome: The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations